ci: move benchmark runners to c7gd.metal (Graviton3) - #9743
ci: move benchmark runners to c7gd.metal (Graviton3)#9743joseph-isaacs wants to merge 4 commits into
Conversation
Switch every benchmark workflow that ran on c6id off x86 and onto c7gd: the label-triggered PR benchmarks (SQL presets, compress, random access, string), the post-merge develop benchmarks, and the nightly SQL matrix. The bench job runs on c7gd.metal and the binary is built on c7gd.8xlarge so `-C target-cpu=native` matches the benchmark host. Both jobs pin the arm64 image since `bench-dedicated` defaults to x64. `sql-bench-matrix.yml` gains `build_machine_type` and `machine_image` inputs alongside `machine_type` so a caller can still pick another architecture consistently; the nightly matrix passes all three. The setup-duckdb action now picks the DuckDB CLI archive by `uname -m` instead of hard-coding amd64. Signed-off-by: "Joe Isaacs" <joe.isaacs@live.co.uk> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FUSFWCFtHuDq5uCqDgGm9i
Merging this PR will regress 3 benchmarks
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | random_i8[0.5] |
67.2 µs | 90.6 µs | -25.87% |
| ❌ | WallTime | words_gather_scalar_avx2[65536] |
8.3 µs | 9.4 µs | -11.8% |
| ❌ | Simulation | allocate_drop_bytes[0] |
466 ns | 520.2 ns | -10.41% |
| ⚡ | WallTime | arrow_checked_add_u32_neon[16384] |
20.5 µs | 12.7 µs | +60.84% |
| ⚡ | Simulation | random_i16[0.95] |
94.4 µs | 76.4 µs | +23.64% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing claude/benchmark-runners-c7gd-metal-3mhrqb (dd8b594) with develop (792cef7)
Footnotes
-
206 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
|
CodSpeed check: the four reported regressions are measurement noise, not this PR's. The diff touches only benchmark workflow YAML and the DuckDB CLI download in Generated by Claude Code |
Polar Signals Profiling ResultsLatest Run
Previous Runs (18)
Powered by Polar Signals Cloud |
Benchmarks: String Encoding 📖Commits: PR vortex / vortex-file-compressed / ms (0.820x ✅, 7↑ 0↓)
vortex / vortex-file-compressed / % (1.000x ➖, 0↑ 0↓)
|
|
String Encoding benchmark failed on
What differs between the green arm64 test job and this failure, in the order I'd check them:
I can't reproduce this from an x86 host, so I'm not pushing a fix. Suggested next step is to run the string bench on an arm64 box once without The compress, random-access and SQL benches are still running; I'll report whether they hit the same thing. Generated by Claude Code |
Benchmarks: Compression 📖Commits: PR vortex / vortex-file-compressed / ns (0.898x ✅, 13↑ 2↓)
vortex / vortex-file-compressed / bytes (1.000x ➖, 0↑ 0↓)
vortex / vortex-file-compressed / ratio (0.967x ➖, 9↑ 6↓)
vortex / parquet / ns (0.944x ➖, 10↑ 1↓)
vortex / parquet / bytes (1.000x ➖, 0↑ 0↓)
vortex / arrow-ipc / ns (0.614x ✅, 23↑ 0↓)
vortex / arrow-ipc / bytes (1.000x ➖, 0↑ 0↓)
|
Benchmarks: PolarSignals Profiling 📖Commits: PR datafusion / vortex-file-compressed / ns (0.636x ✅, 9↑ 0↓)
No file size changes detected. |
|
Compression benchmark failed on That is the same shape of failure as the OnPair one above: the decoded byte count disagrees with the So the working theory is a miscompile or exposed UB under Generated by Claude Code |
Benchmarks: TPC-H SF=1 on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.728x ✅, 22↑ 0↓)
datafusion / vortex-compact / ns (0.699x ✅, 22↑ 0↓)
datafusion / parquet / ns (0.804x ✅, 22↑ 0↓)
duckdb / vortex-file-compressed / ns (0.658x ✅, 22↑ 0↓)
duckdb / vortex-compact / ns (0.678x ✅, 22↑ 0↓)
duckdb / parquet / ns (0.996x ➖, 1↑ 0↓)
File Size Changes (5 files changed, +0.0% overall, 4↑ 1↓)
Totals:
|
Benchmarks: FineWeb NVMe 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.592x ✅, 9↑ 0↓)
datafusion / vortex-compact / ns (0.543x ✅, 9↑ 0↓)
datafusion / parquet / ns (0.630x ✅, 8↑ 0↓)
duckdb / vortex-file-compressed / ns (0.457x ✅, 9↑ 0↓)
duckdb / vortex-compact / ns (0.594x ✅, 9↑ 0↓)
duckdb / parquet / ns (0.809x ✅, 7↑ 0↓)
File Size Changes (1 files changed, +0.0% overall, 1↑ 0↓)
Totals:
|
Benchmarks: TPC-H SF=1 on S3 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.005x ➖, 1↑ 3↓)
datafusion / vortex-compact / ns (1.093x ➖, 0↑ 3↓)
datafusion / parquet / ns (0.933x ➖, 1↑ 0↓)
duckdb / vortex-file-compressed / ns (1.118x ➖, 0↑ 3↓)
duckdb / vortex-compact / ns (1.126x ➖, 0↑ 4↓)
duckdb / parquet / ns (1.027x ➖, 0↑ 0↓)
|
Benchmarks: TPC-DS SF=1 on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.774x ✅, 98↑ 0↓)
datafusion / vortex-compact / ns (0.738x ✅, 99↑ 0↓)
datafusion / parquet / ns (0.782x ✅, 97↑ 0↓)
duckdb / vortex-file-compressed / ns (0.708x ✅, 99↑ 0↓)
duckdb / vortex-compact / ns (0.668x ✅, 98↑ 0↓)
duckdb / parquet / ns (0.878x ✅, 61↑ 0↓)
File Size Changes (3 files changed, -0.0% overall, 1↑ 2↓)
Totals:
|
Benchmarks: FineWeb S3 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.242x ➖, 0↑ 3↓)
datafusion / vortex-compact / ns (1.059x ➖, 2↑ 3↓)
datafusion / parquet / ns (0.906x ➖, 1↑ 0↓)
duckdb / vortex-file-compressed / ns (1.123x ➖, 0↑ 0↓)
duckdb / vortex-compact / ns (1.062x ➖, 0↑ 0↓)
duckdb / parquet / ns (0.227x ✅, 9↑ 0↓)
|
Benchmarks: Clickbench on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.679x ✅, 42↑ 0↓)
datafusion / vortex-compact / ns (0.629x ✅, 42↑ 0↓)
datafusion / parquet / ns (0.676x ✅, 43↑ 0↓)
duckdb / vortex-file-compressed / ns (0.712x ✅, 41↑ 1↓)
duckdb / vortex-compact / ns (0.752x ✅, 38↑ 1↓)
duckdb / parquet / ns (0.745x ✅, 43↑ 0↓)
File Size Changes (83 files changed, -0.0% overall, 41↑ 42↓)
Totals:
|
Benchmarks: Clickbench Sorted on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.797x ✅, 8↑ 1↓)
datafusion / vortex-compact / ns (0.635x ✅, 10↑ 0↓)
datafusion / parquet / ns (0.717x ✅, 10↑ 0↓)
duckdb / vortex-file-compressed / ns (0.716x ✅, 10↑ 0↓)
duckdb / vortex-compact / ns (0.677x ✅, 10↑ 0↓)
duckdb / parquet / ns (0.705x ✅, 10↑ 0↓)
File Size Changes (200 files changed, -0.0% overall, 100↑ 100↓)
Totals:
|
Benchmarks: TPC-H SF=10 on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.662x ✅, 22↑ 0↓)
datafusion / vortex-compact / ns (0.617x ✅, 22↑ 0↓)
datafusion / parquet / ns (0.658x ✅, 22↑ 0↓)
duckdb / vortex-file-compressed / ns (0.665x ✅, 22↑ 0↓)
duckdb / vortex-compact / ns (0.699x ✅, 21↑ 0↓)
duckdb / parquet / ns (0.874x ✅, 12↑ 0↓)
File Size Changes (5 files changed, +0.0% overall, 4↑ 1↓)
Totals:
|
Benchmarks: Statistical and Population Genetics 📖Commits: PR How to read Verdict and Engines
duckdb / vortex-file-compressed / ns (0.747x ✅, 9↑ 0↓)
duckdb / vortex-compact / ns (0.762x ✅, 9↑ 0↓)
duckdb / parquet / ns (0.878x ✅, 7↑ 0↓)
File Size Changes (1 files changed, -0.0% overall, 0↑ 1↓)
Totals:
|
Benchmarks: TPC-H SF=10 on S3 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-compact / ns (1.006x ➖, 2↑ 4↓)
datafusion / parquet / ns (1.010x ➖, 1↑ 0↓)
duckdb / vortex-compact / ns (1.449x ❌, 0↑ 17↓)
duckdb / parquet / ns (1.025x ➖, 0↑ 0↓)
|
Benchmarks: Appian on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-compact / ns (0.686x ✅, 8↑ 0↓)
datafusion / parquet / ns (0.691x ✅, 8↑ 0↓)
duckdb / vortex-compact / ns (0.595x ✅, 8↑ 0↓)
duckdb / parquet / ns (0.644x ✅, 7↑ 0↓)
No file size changes detected. |
Benchmarks: Random Access 📖Commits: PR How to read Verdict and Engines
vortex / arrow-ipc / ns (0.977x ➖, 4↑ 9↓)
random-access / vortex-file-compressed / ns (0.892x ✅, 10↑ 1↓)
random-access / parquet / ns (0.522x ✅, 18↑ 0↓)
random-access / lance / ns (1.052x ➖, 0↑ 5↓)
|
Two examples reproduce the Vortex file read failures seen once the benchmarks moved to c7gd.metal (OnPair/FSST decoded bytes disagree with uncompressed_lengths): * `sve_widening_sum`: dependency-free. With `-C target-cpu=neoverse-v1` (what `target-cpu=native` resolves to on Graviton3), rustc 1.98 / LLVM 22 miscompiles a widening `u8 -> usize` sum at an SVE vector length of 256 bits or more; the same binary is correct at 128 bits, and u16/u32/i32 sums are correct everywhere. Both codecs size their decode buffer with exactly that sum over the file's `u8` string lengths, so the buffer comes out about half the needed size. * `arm64_repro`: the same fault through the real path, on the ClickBench URL column, as five independent stages (in-memory OnPair and FSST, then file write/read with OnPair, FSST and the default compressor), with a `--dump` mode that writes every decoded child and its widening sum so two runs can be diffed. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FUSFWCFtHuDq5uCqDgGm9i
`-C target-cpu=native` on c7gd.metal resolves to neoverse-v1, and rustc 1.98 / LLVM 22 then miscompiles widening `u8 -> usize` sums at 256-bit SVE. OnPair and FSST size their decode buffers with that sum over a file's `u8` string lengths, so every SQL, compression and string benchmark that read a Vortex file failed on the new runners. Add `-C target-feature=-sve,-sve2` to the three benchmark build steps until the compiler is fixed; the standalone repro is `benchmarks/string-bench/examples/sve_widening_sum.rs`. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FUSFWCFtHuDq5uCqDgGm9i
The string-bench library is compiled out without the feature, so the example must declare it as required, like the binary already does; workspace builds without the feature (the musl test job) otherwise fail to compile it. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FUSFWCFtHuDq5uCqDgGm9i
|
going to try #9761 |
Summary
Moves every benchmark workflow that ran on
c6id(x86, Ice Lake) ontoc7gd(arm64, Graviton3 with local NVMe). This covers the label-triggered PR benchmarks (action/bench-sql,action/bench-sql-compact,action/bench-all, compress, random access, string), the post-merge develop benchmarks, and the nightly SQL matrix. The CodSpeed and GPU benchmarks already pick their own instance families and are untouched.Because
c6idandc7gddiffer in CPU architecture, the switch is not a one-word change: the build job has to move too, the runner image has to be spelled out, and the DuckDB CLI download has to become arch-aware. Results from this PR's benchmark runs will be compared against the x86 baseline ondevelop, so the first comparison is expected to be noisy until the develop baseline is regenerated on the new hardware.Changes
pr-bench-runner.ymlanddevelop-bench.yml: bench job runs onc7gd.metal, build job onc7gd.8xlargeso-C target-cpu=nativematches the benchmark host. Both jobs pinimage=ubuntu24-full-arm64-pre-v2sincebench-dedicateddefaults to an x64 image.sql-bench-matrix.yml:machine_typenow defaults toc7gd.metal, and two new inputs,build_machine_type(defaultc7gd.8xlarge) andmachine_image(defaultubuntu24-full-arm64-pre-v2), keep the build job and image consistent with whatever family a caller picks.nightly-bench.yml: matrix entry renamed fromx86/c6id.metaltoarm64/c7gd.metaland passes the new build/image inputs..github/actions/setup-duckdb: selectsduckdb_cli-linux-{amd64,arm64}.zipfromuname -minstead of hard-coding amd64.Checks run:
yamllint --strict -c .yamllint.yamlon every changed file. No Rust changes.🤖 Generated with Claude Code
https://claude.ai/code/session_01FUSFWCFtHuDq5uCqDgGm9i
Generated by Claude Code